Papers with spatial reasoning capabilities
TopViewRS: Vision-Language Models as Top-View Spatial Reasoners (2024.emnlp-main)
Copied to clipboard
| Challenge: | Top-view perspective is a typical way in which humans read and reason over different types of maps, but spatial reasoning capabilities of modern VLMs in this setup remain unattested and underexplored. |
| Approach: | They introduce a top-view spatial reasoning dataset and use it to evaluate VLMs across 4 perception and reasoning tasks with different levels of complexity. |
| Outcome: | The proposed model can understand and reason over spatial relations from the top view and can be controlled at different granularities of spatial reasoning. |
Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models (2026.acl-long)
Copied to clipboard
Jiahuan Zhang, Shunwen Bai, Tianheng Wang, KaiWen Guo, Zijia Song, Hanqing WU, Guozheng Rao, Kai Han, Kaicheng Yu
| Challenge: | Existing benchmarks explore aspects of threedimensional spatial reasoning and visual-language reasoning in dynamic environments, but they are unable to perform well on 3D spatial deformation reasoning. |
| Approach: | They propose to use a ladder competition format to assess the model's spatial deformation reasoning abilities to determine its performance. |
| Outcome: | The proposed framework assesses the performance of Vision-Language Models in spatial deformation reasoning tasks. |
An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large Multimodal Models (LMMs) have shown impressive generalization ability on vision and language tasks, but their spatial understanding is under-explored. |
| Approach: | They construct a VQA dataset to analyze LMMs' spatial reasoning capabilities. |
| Outcome: | The proposed model is stronger at basic object detection than complex spatial reasoning. |
iVISPAR — An Interactive Visual-Spatial Reasoning Benchmark for VLMs (2025.emnlp-main)
Copied to clipboard
| Challenge: | Vision-Language Models (VLMs) struggle with spatial reasoning and visual alignment, despite their performance on 2D tasks. |
| Approach: | They propose a multimodal benchmark to evaluate VLMs' spatial reasoning capabilities based on the sliding tile puzzle . |
| Outcome: | The proposed model performs better on 2D tasks compared to 3D or text-based settings, but struggles with complex spatial configurations and consistently falls short of human performance. |